Papers with evaluation costs
Early-Exit and Instant Confidence Translation Quality Estimation (2026.eacl-long)
Copied to clipboard
| Challenge: | Quality estimation models are often opaque and computationally expensive, making them impractical to be part of large-scale pipelines. |
| Approach: | They propose an uncertainty-aware quality estimation model that matches previous approaches at a fraction of their costs. |
| Outcome: | The proposed method reduces evaluation costs by 50% and improves reranking performance. |
Less is More for Long Document Summary Evaluation by LLMs (2024.eacl-short)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown promising performance in summary evaluation tasks, yet they face challenges such as high computational cost and the Lost-in-the-middle problem where important information in the middle of long documents is often overlooked. |
| Approach: | They propose a novel method which extracts key sentences from a long source document and then evaluates the summary by prompting LLMs. |
| Outcome: | The proposed method significantly reduces evaluation costs and exhibits a higher correlation with human evaluations. |
Mergenetic: a Simple Evolutionary Model Merging Library (2025.acl-demo)
Copied to clipboard
| Challenge: | Recent work shows that combining model merging with evolutionary algorithms can boost performance, but there is currently no library for experimenting with different evolutionary algorithms and merging methods. |
| Approach: | They propose an open-source library for evolutionary model merging that enables easy composition of merging methods and evolutionary algorithms while incorporating lightweight fitness estimators to reduce evaluation costs. |
| Outcome: | The proposed library produces competitive results across languages and tasks using modest hardware. |
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? (2025.acl-industry)
Copied to clipboard
Md Tahmid Rahman Laskar, Mohammed Saidul Islam, Ridwan Mahbub, Ahmed Masry, Mizanur Rahman, Amran Bhuiyan, Mir Tafseer Nayeem, Shafiq Joty, Enamul Hoque, Jimmy Huang
| Challenge: | Large Vision-Language Models (LVLMs) are expensive and time-consuming to evaluate . however, they are limited in their use in industrial settings due to their limited availability and limited resources. |
| Approach: | They evaluate 13 open-source LVLMs as judges for diverse chart comprehension and reasoning tasks. |
| Outcome: | The proposed models can be used to assess chart comprehension and reasoning tasks, but they are expensive and time-consuming. |
MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing LCU benchmarks for large language models often result in prohibitively high evaluation costs . existing benchmarks exhibit significant redundancy, which means inefficiency in evaluation . |
| Approach: | They propose a data compression method tailored for long-text data with sparse information characteristics. |
| Outcome: | The proposed method reduces evaluation costs to 4.5% of the long-text benchmark LongBench . the proposed method is based on a long-term LCU benchmark with sparse information characteristics . |
Evaluating the Creativity of LLMs in Persian Literary Text Generation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Prior research has focused primarily on English, with limited exploration of non-English literary traditions and without standardized methods for assessing creativity. |
| Approach: | They build a dataset of user-generated Persian literary spanning 20 diverse topics and assess model outputs along four creativity dimensions . |
| Outcome: | The proposed models generate Persian literary text enriched with culturally relevant expressions. |
Transfer-Aware Data Selection for Domain Adaptation in Text Retrieval (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to improve domain adaptation do not guarantee improved adaptability, but may negatively impact model performance. |
| Approach: | They propose a framework that can effectively improve model adaptability by selecting beneficial data without evaluating all source data. |
| Outcome: | The proposed framework improves model adaptability by selecting beneficial data without evaluating all source data. |
SubLIME: Subset Selection via Rank Correlation Prediction for Data-Efficient LLM Evaluation (2025.acl-long)
Copied to clipboard
Gayathri Saranathan, Cong Xu, Mahammad Parwez Alam, Tarun Kumar, Martin Foltin, Soon Yee Wong, Suparna Bhattacharya
| Challenge: | Large language models and datasets have made benchmark evaluations computationally prohibitive. |
| Approach: | They propose a framework that reduces evaluation costs by 80% to 99% while preserving ranking fidelity. |
| Outcome: | The proposed evaluation reduces evaluation costs by 80% to 99% while preserving ranking fidelity. |
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities (2025.acl-long)
Copied to clipboard
| Challenge: | ONEBench enables custom benchmarks for specific capabilities while reusing and aggregating samples. |
| Approach: | They propose a new paradigm that consolidates individual evaluation datasets into a unified, ever-expanding sample pool. |
| Outcome: | The proposed model evaluation framework is based on dynamic, sample-level evaluation. |